Diverse Counterfactual Explanations

Summary

DiCE generates multiple different ways to change an input so that a model gives a desired prediction.
It belongs to Interpretable Machine Learning#Counterfactual explanations, and Evolutionary Algorithms can be used as one search engine inside DiCE.

Definition

Diverse Counterfactual Explanations (DiCE) is a method for generating counterfactual explanations for machine learning predictions.

A counterfactual explanation asks:

What small change to this input would change the model prediction?

DiCE extends this idea by asking:

What are several different small changes that could change the model prediction?

Why "diverse" - why diversity matters

One counterfactual can be technically correct but practically unhelpful.

DiCE tries to generate a set of counterfactuals that are:

How it works

  1. Take an original input x and its current prediction.
  2. Define a desired prediction.
  3. Generate candidate counterfactuals x′.
  4. Score each candidate by several criteria:
    • does it produce the desired prediction?
    • how far is it from the original input?
    • how different is it from other counterfactuals?
    • does it satisfy constraints?
  5. Search for a set of counterfactuals that balances these criteria.

Objective

DiCE is usually balancing three goals:

Validity

The counterfactual should flip the prediction.
Example:

Proximity

The counterfactual should be close to the original input.
Example:

Diversity

The counterfactuals should not all say the same thing.
Example:

Search engines for DiCE

DiCE is the explanation method, but it still needs a search engine to find good counterfactuals.

The search engine is the procedure that tries many possible x′ values and looks for candidates that are valid, close, diverse, and realistic.

Important

The search engine affects what kinds of counterfactuals are found. DiCE's explanation quality depends not only on the idea of diversity, but also on whether the search procedure can find realistic and actionable candidates.

Search Engines

Search Engine How it works When to use
Random search generates many random changes around the original input
keeps the candidates that work
- use as a simple baseline
- when the feature space is small
- when the model is black-box
- when gradients are unavailable
- less useful when valid counterfactuals are rare
Genetic algorithm keeps a population of candidate counterfactuals and gradually improves them using Evolutionary Algorithms#Genetic Algorithm (GA) - when the model is black-box
- when the search space is mixed: continuous + categorical + integer features
- when the objective is non-smooth or constrained
- when random search is too inefficient
- can be slower because it needs many model evaluations
Gradient-based search directly optimizes the counterfactual input using gradients - useful for differentiable models, especially neural networks
- when features are mostly continuous
- usually faster than random or genetic search
- less suitable for tree models, categorical-heavy data, or hard constraints
Nearest-neighbor / KD-tree search looks for real examples in the dataset that already have the desired outcome - when realism is important
- when you want counterfactuals close to actual observed data
- useful for tabular data
- less useful if the dataset has few examples with the desired outcome
- useful in very high-dimensional spaces where nearest-neighbor search becomes weak

Choosing a search engine

Situation Better engine
Need a simple baseline Random search
Black-box model Random search or genetic search
Many categorical / integer features Genetic search
Mostly continuous features + differentiable model Gradient-based search
Need realistic examples close to observed data Nearest-neighbor / KD-tree search
Valid counterfactuals are rare Genetic search or gradient-based search
Need fast approximate results Gradient-based search if available; otherwise random search

Comparison with other local explanations

Method Main question
LIME What simple local model explains this prediction?
SHAP How much did each feature contribute to this prediction?
Counterfactual explanation What minimal change would flip this prediction?
DiCE What are multiple different ways to flip this prediction?

Pros & Cons

Pros

Cons

Important caveats

Related Notes